IaaS Platform Data Model Documentation
https://app.diagrams.net/#G1wfJENGlZH2LgvAPtc7uryNN1vSIQOyuy




Agents
Agents on workers are responsible for all tasks. There is no 'work' done by server side processes All 'work' such as creating a ceph volume, checking the amount of capacity in a ceph pool, deleting a container etc are all handled by a worker chosen by the server(TODO - Link to the agent selection process)
Agent configuration options
VM_STORAGE_LOCATION - The agent can specify the disk \ volume \ folder that is used to store VM data, this is the disk that used to report available capacity back to the server WORKER_ID

Ceph
TODO - Can this all be sent by the server to the worker at the time of the task allocation? CEPH_CONFIG_FILE CEPH_USERNAME CEPH_KEY

Deploy a new VM request (VM_LAUNCH)
A VM launch payload contains a number of attributes. This describes the customer facing API payload that needs to be recived by the API to process the new VM launch request
Mandatory attributes
Name
RAM(GB)
CPU_Cores
Storage volume dict source_type - from_image, from_volume, blank_volume source_name - name\location of the source volume, not valid if source_type=blank_volume dest_size - Size in GB of the new volume dest_type - ceph or local
Optional
CPU_Type - #TODO - Define a CPU type api endpoint
Description
Array of fixed resources
name
qty
Array of networks to bind too
network.id
Array of labels - an array of key:value pairs or just keys with no value
VDC_ID - The ID of the user's VDC, if none is supplied the system will check for VDC's assigned to the user, if more than 1 exists, the server use the default_vdc attribute on the user account
A number of 'defaults' are applied including storage driver, graphics driver, virtual cpu type etc
The API rx's this launch request and runs a number of checks
Permissions \ ACL check
Does the user have VM_Launch permissions?
Placement - [See here]({{< relref "placement" >}})
Does the user have VM Launch permissions(VM_LAUNCH)
The server breaks down the process of buildign the VM into seperate jobs
Spawn the VM with the requiste volumes attached and the network port Once VM is spawned - Fetch the VM id and VNC port info, report it back to the API
Key points
Crazy simple
The architecture for this project is designed to be as simple as possible. It consists of a front end, a back end, and a database, keeping it straightforward and easy to understand.
Monolithic
{ProjectName} is structured as a monolith, avoiding the complexity of multiple microservices or individual projects. It operates as a single cohesive unit, featuring a single API endpoint and a worker agent on every node that handles all tasks.
Modularity in code
For things like RBAC, Quota, Placement decisions etc. These functions will all be called form inside a larger routine, but the function that is being called should be a fairly 'modular' piece of code that can be referenced by other functions elsewhere, it can be upgraded independant of other pieces of code and should be held in a seperate file, possibly a seperate repo?
Master server
There is always a designated master server. The master is decided using the RAFT protocol The master server is responsible for
Scheduling tasks - The master server runs the 'cron' scheduler that assigns the periodic tasks to the agents
When a server transitions from SLAVE to MASTER state an entry is logged
Scheduled tasks
There are a handful of tasks that need to be run routinely, these will be handled by the master server
Agent heartbeat timeout
The master server runs a simple SQL query searching for hosts that have last_heartbeat_time > 30 seconds and updates the status column to 'offline' then rund the handler routine for 'AgentHeartbeatMissed'



Project Invites
Users can be invited into projects
Invites can (optionally) specify a team for the user to join
Invite characteristics:
Expiration date
Status indicating acceptance
Track the user issuing the invite
Invite rules:
Users can only invite others if they have 'SEND_INVITE' permission
Users can only add a team to the invite if they are a member of that team

Permissions Model
Core Permission Principles
DENY permissions override ALLOW permissions
No specific ordering of permissions
If no DENY is found and an ALLOW is found, the user has access
Permissions collection process:
Collect all relevant permissions
Where user/group AND object ID/type are mentioned
Sort DENY permissions first
Permission Scope
Permissions can be applied at:
Project level
VDC level
Individual object level
Permissions filter down hierarchically
Example
If a user has 'create VM' permission on a project:
They will have this permission in all VDCs
A specific VDC can have a DENY for 'create VM' to restrict access
Permitted Actions
Actions are specific requests a user can make against an object:
Hard-coded and predefined
Granted as ALLOW (1) or DENY (0)
No generic 'read', 'write', or 'edit' concepts
Actions define specific capabilities
Sample Actions
VM_LAUNCH: Launch new VMs in a VDC
VM_LIST: List all VMs in a VDC
VM_DELETE: Delete VMs in a VDC
SEND_INVITE: Send invites for a VDC
OWNER: Indicates that the user has the admin permissions to make changes to the object(Used on a VDC and a project)
VNC_CONSOLE - Permits access to the Virtual Macine VNC console
CONTAINER_CONSOLE - Permits access to the interactive terminal for a docker container
Teams
Logical grouping of multiple users for permission management
Users can be team members, but team membership is optional
Teams are limited to a single project
Users cannot be in a team if they are not project members


Data models
Base model
All models in the system share a set of base attributes:
Unique ID - UUID
Created date - default to now()
Last modified date - required- default to now() when created
Deleted date - default null
Deleted - boolean default 0\false
Visible - Boolean default true
Name
Description
Status
Created by - id of the user who created the object
—-------------
Universe
Top level of the hieirachy, everything ultimatley belongs inside a region
—-------------

Project
All Virtual Datacenters (VDCs) exist within a project
Permissions can be applied at a project level
Users are members of a project via a many-to-many UserMembership model
Membership does not imply permissions
Users must be granted specific permissions to perform actions
Every project has a default 'admin' team with full permissions on all objects
When a user creates a project they are assigned the OWNER permission to that object ID
Properties
Universe - fk link to a universe_id
—-------------


Virtual Datacenter (VDC)
Exists within a specific region
All workloads scheduled within a VDC are restricted to its region
When a user creates a VDC they are assigned the OWNER permission to that object ID
—-------------


Region
Regions can exist beyond the scope of a Universe, becuase a universe is typically an instance of a cloud, but a region can exist within 2 clouds
Properties
Private region - Indicates if the region is ‘public’ or available to eveyrone required, default false
Country - Required
Advertised address  -Eg ‘Washingthon State, USA’  -optional
Region abbreviation - Short text based name - Required - max len 16


—-------------

Region access
Many to many table granting access to a region
Access is based on a user being a member of a project

Properties
project _id - fk to project - required
Region_id - fk to a region - required

—-------------

Labels
Super flexable, any label can be applied to almost anything
Properties
Target object type(Will be a model name, i.e workload_host, volume, vm, container etc) - required
Target object id - the ID of the object this label applies too - required
Owner id - required
Label key - required
Label value - optional
Admin_only_view - indicates if the label is only avialbale fo admins to view - optional, default false
Admin_only_set - - indicates if the label is only avialbale fo admins to edit, if false, anyone with permission to the object it’s attached too can edit the label - optional, default false
—-------------


Worker Agent
1:1 assignment to a workload host
Configurable attributes:
Workload_host-id - required - FK to a workload_host id
available_for_scheduling: Boolean flag to control workload flow
placement_priority: Integer, defaults to 100
heartbeat_status: Represents recent host communication
last_heartbeat_time: Timestamp of last heartbeat
status: Can be 'available' or 'offline'
Agents and hosts have been combined, there was a 1:1 mapping anyway, so pointless


Workload Host
Can be a physical server, VM, or device capable of running workloads
Exists within a ‘region’
Properties
Region id - required
System manufacturer - optional
System model - optionla
System physical identifier(If known by a different name in the DC, could be rack, RU description etc) optional
DCIM identifier - could be an id or a url to netbox etc. Optional
Installed date - Date of installation in the region, optional
Installed status(Used for adding machines when their status is planned and this is the planned install date) - optional
available_for_scheduling: Boolean flag to control workload flow
placement_priority: Integer, defaults to 100
heartbeat_status: Represents recent host communication
last_heartbeat_time: Timestamp of last heartbeat
status: Can be 'available' or 'offline'

—-------------

Workload Host Fixed Resources
Tracks individual resources like GPU, NVME, NIC typically used in PCI passthrough mode
In a machine with 8 GPU’s, we will have 8 device entries here
Attributes:
Status (in use or not)
Type (GPU, NIC, NVME, PCI, Other)
Resource model
Resource manufacturer
Physical address - mac address, pci addres, device number etc
Linked to a worker agent- fk to worker_agent_id


—-------------

Workload Host Pooled Resources
Tracks aggregated resources like CPU, RAM, GPU, network bandwidth
Attributes:
Total quantity - required - number, could have a decimal, default 0
Quantity in use - required - number, could have a decimal, default 0
Quantity available - required - number, could have a decimal, default 0
Worker_agent id - fk to worker_agent_id
Resource type - string, required


—-------------
Workload host OVS Bridge
Report back the OVS bridges that are available on the host. These are typically used to attach containers or VM’s directly to a physiacal network for internet access
This is a simple mapping of OVS bridge name to workload_host_id
This table\model can be used by the placement routine to filter hosts with a specific network availabel to them

—

—-
Workload
A generic model used by all workload types
Current host - fk to workload_host_id
Type (Container or VM,or NSController) - required
VDC ID - fk to vpc_id
Launch attributes
Status
Container_id


Namespace controllers
Tracks the docker namepace controllers
This lets us assign IP’s to the controllers for both LAN(VXLAN tunnels) as well as when attaching them to public\physical networks like the internet
If we deploy NS controllers in a central region for the purposes of doing a global proxy this will be usefull, becuase there wont necessarily be an associated workload
We also need to track the nscontrollers when we are cooridinating sit to site VPN’s


—-------------

Network
A network could be a virtual network like a vxlan
Or it could be a public network, shared across all VPC’s in a region
A public network is iusaully goig to be a physical network on a host or set of hosts that is used to conect to the internet or a real world network
A physical nework could be attached to NScontrollers or VM’s
When a VM launch is requested, placement will need to take into consieration any physical network requirements to check if there are any hosts that can serve the workload and have this network attached
The idea with a physicla network is for docker hosts that most ns controleers will attach to a physical network(OVS Bridge)  to allow their containers to get internet access outbound

VDC id - fk link to vdc_id
IPv4 CIDR (optional, can generate random /24 in 10.x range)
IPv4 gateway IP
IPv4 DNS servers
IPv6 CIDR
IPv6 gateway IP
VNI -   VXLAN id - required
OVS Bridge name
—-------------


Network Port
Properties
Linked to a Network - required - fk to network_id
Linked to a Workload - required - fk to a workload_id
IP address
MAC Address
—-------------

Port_ACL
Properties
Network port - fk to network_port_id
Port - TCP or UDP port # - optional
Protocol type - TCP\UDP\ICMP etc - optional
Action - Usually block or allow - optional

—-------------

Image
Properties
Location - required
Location type - text - Could be used to identify images stored in S3 ns NFS etc
Size - required
Operating system family optional
Operating system version - optional
—-------------


Volume
Type (Ceph or local) - required
Path (RBD location or local filename) - required
Size in GB - should be a number, can have decimal places
Ceph pool id - optional fk to ceph_pool_id
Image id - optional
VPC - fk to a vpc_id - required
—-------------
Volume Attachment
A volume can be attached to multiple workloads
Volume ID - fk to volume id -required
Workload ID - fk to workload id  - required
Attachment type (SCSI, CD/ISO, Sata, VirtIO, etc.)
Status - Planned_attach, attached, planned_detach, detached
—-------------
Resource Consumption
Tracks resource usage across workloads
Resource ID
Resource type (Fixed or dynamic)
Workload ID
—-------------
Ceph Pool
FSID - required
Pool name - required
Number of volumes - int - default 0
Total pool size in GB - requried - default float 0
Free space in GB - requried - default float 0
Last updated - required - date time - default now()
—-------------
User
Properties
First name
Last name
Email address
OIDC ID
Development challenges
Message queuing
Status - Working POC produced 
TODO - Define job payload format including 'operations' and 'desired_states'
Workflow management
Define the workflows for each type of request
Ceph volume management

VM Management
Create a VM VM Operations - Start\Stop\Restart
Container Management
Create a VM VM Operations - Start\Stop\Restart
Network ops
Create a VXLAN Network Apply ACL's to a specific vNIC
API Creation
Define the API routes Construct ACL mechanisim Define data models

DNS
There will be an authroitate DNS server running somewhere that will serve real world DNS Users can create their own DNS domains, providing a DNSaaS type function and VM's will be registered against the primary domain, offering name resolution for VM's
+++ archetype = "chapter" title = "Introduction" weight = 1 keywords = ["{ProjectName}", "OpenStack", "virtual machines", "system architecture", "use cases"] +++

Overview
What is {ProjectName}?
{ProjectName} is a specialized project designed to streamline IaaS Operations. While it aims to replace OpenStack in these particular use cases, it does not intend to offer a comprehensive one-to-one replacement for OpenStack.
Project Focus
{ProjectName} prioritizes the efficient scheduling of workloads(VM’s, Containers or clusters) and their supporting infrastructure, catering to a select set of use cases. Unlike many other systems, {ProjectName} adopts a highly opinionated stance on system architecture, tailored specifically to a given configuration of environment, network architecture, and storage system layout.
Key deliverables
Build and manage virtual machines and docker\podman containers
Multi-tenant native system
Support the use of local volumes, parallel filesystems and Ceph volume storage
Support PCI device passthrough for use cases like SR-IOV, GPU passthrough, Other PCI card passthrough
Support PCI device affinity to enable NVLink aware VM Scheduling - [See here]({{< relref "placement" >}})
Granular Access Controls \ RBAC
Comprehensive auditing, everything is audited and logged with a target for compliance
Target Audience
{ProjectName} is not a universal solution. It's intentionally designed to serve a niche audience with clear objectives. By focusing solely on the task of building and managing virtual machines and containers, {ProjectName} avoids unnecessary complexity and bloat often found in other systems.
{{% notice note %}} Please note that {ProjectName}'s design and functionality are optimized for specific environments and use cases. It may not be suitable for all scenarios. {{% /notice %}}
The keyv target audience are IaaS operators looking to deliver hardware on demand to mutiple customers, specifically customers looking for on-demand, self service, secure and performant GPU accellerated VM’s or containers
This project does not replace a Kubernetes or Slurm cluster. If a customer requires a single large cluster with a single tenancy and little churn this this project adds little value.
Limitations
Initial project will be 'single region' only, the idea being if a second region is required you will simply launch a second instance of the system.

BYO Cloud
A customer can enroll their own resource(Bare metal or VM) in their own datacenter to {ProjectName} and enable is as a ‘Compute Host’, this will let customer sue a single pane of glasss for their cloud and on prem operations. 
Having these resources bound together in a single API with inter site comms is a huge USP that will be highly regarded by engineers and developers.
A customer could use this concept to bring their local dev machine into the cloud for test or dev purposes or they could enroll their AWS\Azure\GCP resources into {ProjectName} and manage them in the exact same way they manage their cloud compute
A customer could have a dedicated ‘region’ that is available only to themselves to enable them to achieve compliance or data soverignty requirements
The customer could also offer this capacity to the global\public pool, enabling them to monetise underutilised compute power from the same portal as their cloud consumption.
Concepts
Heirachichal model
Universe - Top level model used for total isolation of clouds - Could be used at the white label level?
Project - Contains VDC's, users have an association with one or more Projects (Could be seen as a ‘company’ or a ‘tenancy’)
VDC - Belongs to a Team, is specific to a region 
Resource - Exists within a VDC
In openstack they have 'projects' here we have VDC's(Virtual Datacenters). A VDC is a virtual representation of what a customer would expect in a typcial colocation setup, it includes their private network, their router\firewall device, some shared storage and some servers(Could be VM's or containers)
Many users can have access to a given VDC
When a user is created a blank VDC is created for them, no objects are created on the physical infra untill a service is provisioned
A user can invite another user to join their VDC One user can be a member of many VDC's the permission model is used to control which permissions user have in a given VDC
VM's dont go into error when launched
Orphans are impossible
Stock reporting is included out of the box
Billing data is exposed for integration into any app
Designed for real time integration, no 'shadow' DB's are required
Fine grained ACL's to enable easy integration
User scoped tokens to allow an intermediary application to use properly scoped tokens to avoid admin bugs and exploits
All actions are logged with comments, so if a descrutive(Or any) action is taken we can know when, where, by whom and why
SSH access to VM's can be enabled at creation time
SSH access can be via intergrated SSH CA for SSO or via SSH keys or password
Post VM deployment validation process occurs to ensure the VM is deployed as expected
'Jobs' cant get stuck in a queue causing all other operations to fail
Consolidation (Stack\Affinity): Places instances together for performance benefits.
Dispersion (Spread\Anti-Affinity): Distributes instances for fault tolerance.
Visual UI to show VM placement and results of stack or spread
Organiser to do VM live migrations to tidy up VM's or do maintenance
Competitive Analysis
AWS Management Console: Feature-rich, but can be overwhelming for new users
Google Cloud Platform Console: Modern UI, good usability, but sometimes slow response times
Azure Portal: Highly customizable, but complex navigation
DigitalOcean Dashboard: Simple and intuitive, but lacks some advanced features
Linode Cloud Manager: Good performance, but UI feels outdated
VMware vSphere Client: Powerful, but requires significant learning curve, on prem only
OpenStack Horizon: Open source, customizable, but inconsistent UI experience
Logging from all containers and workers running on all systems will be directed back to the central API
There is a /log API endpoint that will recieve either a single log entry or a batch of entries
Log data will be stored in Elsaticsearch(Or an equivalent)
Each log payload will be formatted as follows
host_name - The name of the host submitting the data as identified by the system hostname
date_time - The time the log entry was created(Not the time it was received)
SourceIP - The IP of the agent submitting the log entry, this is the source IP of the agent(Usually the public IP)
Log data - This is an array or dict of data including but not limtied to
message
source - Syslog, docker container
severity
container
+++ archetype = "chapter" title = "Message Flow" weight = 3 keywords = ["message flow", "controller", "server", "task", "dispatching", "workers"] +++
Overview
There are one or more API servers in the system fronted by a proxy that allows for sticky websocket connections
Clients connect to a single server in the server 'pool' The server that recieves the conenction subscribes to the task ques for that client ID In the event of a server failure, the clients will reconnect to another server in the pool and that server will subscribe that that clients task queue When a message is publised into a clients task queue the server holdign the websocket connection will recieve the task via the redis subscription and dispatch it to the server
For 'job's the server thats assigning the 'job' to the worker will post the job into the DB, it will thenget he UID of the job and post that UID into the redis task que for the given worker ID. This will allow the responsible server to detect this job ID, query the DB then dispatch the task to the worker
Any server can create a job and assign it to any worker

Networking
All VM and container level networking connectivity will be provided via OVN Each region will have require OVN Controller(Or a cluster) Containers, VM's, Loadbalancers etc will all be connected to the same OVN network in the given VDC
Because we are aiming to keep deployment complexity to an absolute minimum we will not use OVN. A site could consist of a single laptop running docker, it’s complex and overkill to deploy OVN
Instead, we will use OpenVSwitch and OpenFlow to build our own SDN of sorts
VXLAN(Perhaps Geneve) will be used to tunnel traffic between hosts in the same site
In a multi-site configuration, the NSController containers will route traffic between sites. This will happen inside the tenancies and thus has little\no impact on the under;uying infrra
We should prevet\prohibit intersite L2 connectivity and have some sort of site to site router
DHCP, we need to decide how to handle DHCP and popssibly metadata access for VM’s
Can Openflow become a DHCP server? Might remove the need for dhcp container(s) Or possibly we use the ‘controller’ option on the flows to divert DHCP requests to the local worker agent which has the details for the IP\MAC combo of it’s local VM’s


Network Security
Openflow will control the security on the network ports
Container networking \ Container pods
We will need to use and manipulate container namespaces to achieve the level of control required Interesting link https://arthurchiao.art/blog/ovs-deep-dive-6-internal-port/
To effectivly manipulate docker container networking we use a network consider sidecar approach(Concept developed by Cory) called a NamespaceController
In this approach we launch a namespace container first, this container runs a custom image that serves exclusively to manage networking for every container for that tenancy on that host.
This container has an interface added to it thats a member of the OVS bridge, this interface has the openflow rules applied.
Any containers that are launched on this host for this tenancy use the docker –namespace option to direct docker to use the network namespace from the sidecar container, meaning every container shares the same IP.
There is a possible issue here in that if we have 2 containers looking to listen on the same port that one will fail to start, in this scenario we need muliple namespace instances per host(Do we call these pods?)

docker run -d --name namespace-controller busybox sleep infinite
sandboxKey=$(docker container inspect namespace-controller --format '{{  .NetworkSettings.SandboxKey  }}')
echo $sandboxKey
sudo ln -s $sandboxKey /run/netns/ns-docker-namespace-controller
ip netns




ovs-vsctl add-port br-int port-1 -- set Interface port-1 type=internal


# 3. Create a Linux namespace (if not already exists)
ip netns add ns-docker-namespace-controller


# 4. Move the internal port into the Linux namespace
ip link set port-1 netns ns-docker-namespace-controller


# 5. Bring up the interface inside the namespace
ip netns exec ns-docker-namespace-controller ip link set dev port-1 up


# 6. Assign an IP address (optional)
ip netns exec ns-docker-namespace-controller ip addr add 10.0.0.1/24 dev port-1


# 7. Enable loopback inside the namespace
ip netns exec ns-docker-namespace-controller ip link set lo up


# Verify
ip netns exec ns-docker-namespace-controller ip addr show


First class IPv6 support
(Future)
Seamless integration with talescale \ Netbird \ Wireguard \ Zerotier
(Future) Somehow include the ability for new VM's and containers to join a customers tailscale net seamleslessly https://docs.netbird.io/ - Possible option?
Could also use Slack’s Nebula project but this has some interesting UDP port requirements(Notably UDP port randomisation must be disabled)
This should happen inside the NS controller(BAsically the same thing as an Openstack L3 agent)
WAN connectivity
Connectivity to the internet from VM's can be done by a number of mechanisms, ideas so far include
Mikrotik to perfrom 1:1 NAT - Using a hardware device to perfom the public IP to private IP mapping and using it's API to make intergration simple. This device could also handle the firewall requirement.
Place VM's directly on the WAN network, no NAT or routing required
Have one or more Router VM(s) on the WAN network and it handles all the routing config for the customer
Use the Docker namespace controller, co-located on the same host as the VM, this Container can handle all NAT\firewall \ security \ load balancing etc
All outbound traffic will go via this router, it will have the default GW IP for the network(Every host with a VM on it will have a NSController with a container with the default GW)
Inbound traffic will come into the NSController\Router with the respective public IP
We could achieve load balancing by using BGP from the NS controller!
Similar concept to Openstack’s L3 agent but it’s on every host(Similar to DVR?)
This NScontroller\gateway is a docker container.
For lab\home setups the gateway will use the docker host bridge
We use standard container port forwarding to expose services to the local network.
For enterprise nets we dont use the docker bridge, the ns controller will be launched with no networking then we manipulate the container namespace to add the tenant lan and the WAN network to the container. 
When the container is directly connected to the WAN network, dockers firewalls and port forwarding will not apply
WAN connectivity for containers
For outbound WAN traffic the simplest option is to have the NS controller do NAT via the host, this is ‘standard’ docker networking.
Inbound trafffic for web applications can coem through a proxy
SSH access in can match runpod - Some sort of proxy and runnig gotty
Can we get a native SSH access in?

Containers can have upto 3 network options
Option 1 - Host bridge(Or a specified ‘other’ network)
This will be the default route for the Namespace controller and the container
In a home environment, this will be the simple docker bridge
In a secure environment it will be some ‘other’ network that will give the container routed internet access, probably via NAT
Option 2 - The Region wide SDN
This is our managed network using OVS and VXLAN 
Requires OVS
Option 3 - The inter-region network
Probably a Zero tier interface whereby we can access services on other sites
The encrypted ZeroTier traffic will follow the default route


MAC Address handling
MAC Addresses should be guaranteed to be unique to a VDC When creating a new network port, create a random MAC then do a SQL query for all network_port objects in the DB with a matching MAC and VDC ID. If a match is found then run the function again. Loop 10 times before raising an exception
OVN Free concept
If we use simple OVS and populate the FDB on each node(As neutron does) we can achive a much more simple architecture, no OVN reuired.
https://docs.nvidia.com/networking/display/mlnxofedv522240/sr-iov+live+migration https://arista.my.site.com/AristaCommunity/s/question/0D52I00007ERpvxSAD/veos-and-openvswitch https://arista.my.site.com/AristaCommunity/s/article/vxlan-without-controller-for-network-virtualization-with-arista-physical-vteps https://cwiki.apache.org/confluence/display/CLOUDSTACK/OVS+distributed+routing+and+network+ACL

OVS Rules reference
https://github.com/huangyingting/glb-demo/blob/b398fcff9d29bf45becda1fa755ad522d738f752/ubuntu-ovs.yml#L83

Provisioning flow - New workload
Orchestrator sends new VM\container event to the host
Host creates the OVS port(s)
(Other VM prep operations like storage, hugepages, shared filesystems etc)
VM Boots
Host replies - VM booted!
Orchestrator advises all other hosts with workloads on the same VNI in the same region that a new network member is present by sending the full list of all devices on all hosts to every host, allws the hosts to cross check their flow tables, purge any old dead results and update for the new network member




Permissions
List of permission plags
VM_LAUNCH (VDC)
Permits a user to launch new VM's in the given VDC
SEND_INVITE (VDC)
Permits a user to send invites for the given VDC

Placement
Concepts
A host can be reserved for one or more user\projects. This will allow admins to have customer owned hardware in a shared cluster


Placement process
Lets do as much of this in python code as possible, will allow for better debugging and less load\requirement on SQL server {{% notice alert %}} This process is not complete, these are just thoughts fornow, this flow needs to be extensivly mapped out {{% /notice %}}
Fetch all hosts where available_for_scheduling = True and their heartbeat_status='Alive'
Check if there are any launch_attributes specified if yes then fill the eligable_hosts list with the subset of hosts who have matching attributes
Filter hosts that have available CPU and RAM grerater than the size specified in the launch request
If the VM is using a local disk, filter on avaialble disk capacity
Sort hosts by AllocationPriority (Hosts with the highest prioirty ranked first)
Loop through all resources in the vm request and filter for hosts with that requirement(E.G )
Make a stack vs spread decision, check ENV property(placement_stack_or_spread), if ENV is set to stack then sort by nodes with lowest {resource} count available but with enough to satisfy this request. where {resource} = CPU, GPU, RAM or some other resource as specified from an ENV property placement_stack_resource_type
From the remaining list of hosts, select all hosts with a matching AllocationPriority E.G if there are 5 hosts with AllocationPriority of 100,100,80,50,50. Then select the 2 hosts with AllocationPriority 100
Pick the first host from the random list. If the list is empty return an error 'No capacity to launch this resource'
Hardware device affinity
Aimed at ensuring we can passthrough NVLink GPU pairs as well as GPU\NIC\NVME combos. This feature enables the admin to store a 'map' of PCI devices that shuould be assinged together or NOT be assigned together. The primary use case will be to define a list of GPU pairs that are connected via NVLink and report this data back to the controller The controller will make placement decisions with NVLink allocation in mine. E.G if a VM is scheduled to a host with NVLink enabled and the request is for 2 GPU's the controller will allocate the 2 GPU's that are in an NVlink pair on that host and when the job is sent to the given worker it will define the exact PCI device ID's that should be mapped to the VM

Workflows
Common attributes
All jobs dispatched to workers come with the following params
not_after - date time - Maximum time at which the job can be executed, if the time now is after this time, report back to the server that the job failed and give a reason 'time exceeds not_after' set to a default of now() + 5 mins
worker_id - the id of the worker, A worker doesnt necessarily know it's worker ID, but the server does, it's included here for future use and debugging purposes
job_type - CEPH_CHECK, VM_LAUNCH, VM_STOP, VM_RESIZE
Deploy a VM Job
When dispatching this task to the worker, this should come through as a single multi-stage task. E.G the tasks to create the volume, network and port should all be supplied in the task payload. The MAC address and port names etc are all pre-determinted
Build the storage volumes
If the new volume is to be based on an image then create the volume from an image
The the VM is booting from an ISO then attach the ceph volume as a CDROM device
If the user has requested multiple volume attachments in the create VM request, add a job for each of them
Build the network ports
Does the network exist in OVN,if no, create it
ovn-nbctl lswitch-add {OVN_NETWORK_ID}
Create the port on on the network
ovn-nbctl lport-add {OVN_NETWORK_ID} {OVN_PORT_ID}
ovn-nbctl lport-set-addresses {OVN_PORT_ID} 00:00:00:00:00:01
ovn-nbctl lport-set-port-security {OVN_PORT_ID} 00:00:00:00:00:01
ovs-vsctl add-port br-int {OVN_PORT_ID} -- set Interface {OVN_PORT_ID} external_ids:iface-id={OVN_PORT_ID}
write the supplied XML file to disk
Build the VM
virsh define
Start the VM
virsh start
Get the VNC port #
virsh ....
Respond back to the API with the VNC port # and libvirt VM ID
Agent starting up
When an agent starts on a host it needs to go through quite a few checks then report into the API
Heartbeat (HEARTBEAT)
Agents should regularly heartbeat to the controller so they dont get marked as unavailable This happens on a a timer schedule on the worker, every 10 seconds Payload sent to /jobs/complete TODO: Define the payload of a heartbeat
Creating a batch of VM's
Possibly one of the more complex tasks? A user requests 500 VM's to be scheuled. Clearly the workers are going to do this in parallel, there will need to be some sort of serialisation happening, meaning we need a queue of sorts
Job comes in from API RX'ing server ACK's the request to spawn 500 VM's Creates the requisite tasks in the 'queue', these tasks will include For the batch
Ensure the networks exist(OVN)
Create an OVN port?
For each VM
Check for placement issues(Run placement workflow) and decide on the location of the VM
Create a 'job' for the relevant host to spawn a VM including creating any required local or network disks
'Jobs'
All jobs are outcome focused and deliver a desired state,not a set of instructions. The idea being the same job could be run 5 times but it will only make changes once(Similar to ansible or terraform)
Jobs wil often have multiple parts(E.G create a volume then a VM), if any part of this fails the whole task will fail and the worker will do it’s best to roll back all the other tasks in that job. If it created a volume it’ll delete the volume, care should be taken that if the volume already exissted then it wont delete the volume!
This focus on desired state puts alot of the operational logic into the worker instead of the server, the server simply tasks a worker with ensuring a specific state and the worker is responsible for ensuring the desired outcome
Workers will have tags which is a mechanisim to filter which jobs they can accept, some workers may or may not be able to perform operations on a ceph cluster or host virtual routers, while other systems may host only VM’s and not containers or vice versa
Jobs can also be filtered based on the workers resource availability such as CPU, RAM, HDD or PCI device availability
Jobs are tasks that given workers need to execute, they are assigned to specific workers at the time of the job creation(So it's essental to have up-to-date info about the online\offline workers) 
Jobs are distributed to the workers in serial when the worker checks in and asks for a job bu hitting the /jobs/check endpoint at which time the server will check the DB and if there are any pending jobs for that worker and it will assign the highest priorty task to the node and mark the job as pending Once the task is complete, the worker will report the task status as 'complete' and any relevant details to the /jobs/update endpoint Should the task fail, the worker will set the status to failed via the aforementioned endpoint and it will continue it's main loop, which will fetch anoter task at the next interval If the server detects that the last X jobs have failed for a given worker, it will put the worker into an error state and will refuse to dispatch any subsequent tasks to that worker

Filtering
Proposition based filtering
User provides a set of parameters that the workload requires, which are converted into a set of filters.
E.G if the user requests 4 CPU and 2GB ram thats converted into a filtler which sdearches for a host with that configuration.
The user can add optional filters like - 
Preferred anti-affinity or mandatory anti-afinity {affinity group}
Preferred project group - This could be used in cases where a ‘contract’ is in place and customers workloads are required to goto specific hardware
The filtering logs are included in the details for the workload, so the admin or user can debug the placement and filtering decision.
Order  is important, some filters will exclude hosts that dont fit the criteria, while other filters will order the hosts in list of preference. E.G ‘allocation_priority’ will sort the list of hosts in order of the preferred placement priority, so attention should be paid to the order of hosts.

The filtering process can also handle ‘interruptable’ instance types, if the filter finds 0 available hosts it will re-run the filtering process with ‘interruptable’ flag set during this run workloads that are ‘interruptable’ are excluded from the consideration of host availability and at the end of the run if a host is found and it has an interruptable workload on it, this workload will be destroyed in favor of the incoming workload.
Interruptable instance types can also have priorities set, allowing for multiple tiers of interruptible. Values range from 0-100 where 100  is most important and will be terminated last, will 0 is least important and will be terminated first.
Filtering does not discriminate between VM’s or containers(Or potentially ‘bare metal’) so it’s possible a VM could be deleted in favor of a container or vice versa


Rebuild in place (VM_RESIZE)
Sometimes a VM gets a little cooked or broken and we want to simply delete the definition and rebuild it This process should do a virsh destroy for the VM in question, then rebuild the definition XML file, store that file somewhere sensible then do a virsh define This process wont delete the VM data on disk This process might be used to make hardware changes to the VM that cant be done 'hot'
Ceph stats check (CEPH_CHECK)
Server issues a task to a worker to check the stats of a ceph pool Supplied params
fsid
pool name
The worker checks the totalsize, available space and the number of RBD volumes in the given pool then reports it back to the server via the /jobs/update endpoint

VM Creation
Sample XML

<domain type="kvm">
  <name>linux2022</name>
  <memory unit="KiB">4194304</memory>
  <currentMemory unit="KiB">4194304</currentMemory>
  <vcpu placement="static">2</vcpu>
  <os>
    <type arch="x86_64" machine="pc-q35-6.2">hvm</type>
    <boot dev="hd"/>
  </os>
  <features>
    <acpi/>
    <apic/>
    <vmport state="off"/>
  </features>
  <cpu mode="host-passthrough" check="none" migratable="on"/>
  <clock offset="utc">
    <timer name="rtc" tickpolicy="catchup"/>
    <timer name="pit" tickpolicy="delay"/>
    <timer name="hpet" present="no"/>
  </clock>
  <on_poweroff>destroy</on_poweroff>
  <on_reboot>restart</on_reboot>
  <on_crash>destroy</on_crash>
  <pm>
    <suspend-to-mem enabled="no"/>
    <suspend-to-disk enabled="no"/>
  </pm>
  <devices>
    <emulator>/usr/bin/qemu-system-x86_64</emulator>
    <controller type="usb" index="0" model="qemu-xhci" ports="15">
      <address type="pci" domain="0x0000" bus="0x02" slot="0x00" function="0x0"/>
    </controller>
    <controller type="sata" index="0">
      <address type="pci" domain="0x0000" bus="0x00" slot="0x1f" function="0x2"/>
    </controller>
    <interface type="direct">
  <mac address="00:00:00:04:00:06"/>
  <source dev="desktop-tap1" mode="passthrough"/>
  <model type="virtio"/>
  <alias name="net0"/>
</interface>
    <input type="mouse" bus="ps2"/>
    <input type="keyboard" bus="ps2"/>
    <audio id="1" type="spice"/>
    <video>
      <model type="virtio" heads="1" primary="yes"/>
      <address type="pci" domain="0x0000" bus="0x00" slot="0x01" function="0x0"/>
    </video>
    <memballoon model="virtio">
      <address type="pci" domain="0x0000" bus="0x04" slot="0x00" function="0x0"/>
    </memballoon>
  </devices>
</domain>





Use cases \ user stories
I want to run a docker container, I dont care where it is, I need it to be always available so if it stops, restart it, if the host crashes then respawan it. I need this container to have  a specific type of GPU attached and a specific amount of RAM and CPU cores. I'd like to be able to access it from the internet but possibly only from certain IP's. Ideally I'd like to have a CDN sit in front of this container to serve the content, this CDN will help me with global availability, DDOS protection, stastics and authentication
I want to run a docker container,but I dont have a published container on a public repo, let me provide the Dockerfile and you run it for me, making the image available to me in a private repo. I'll pay for the data hostign costs assuming they are low enough.
I'd like to be able to use docker compose but have the containers launch on the cloud. Kinda like terraform but easier
I'm cory- I need lots of shit running in multiple DC's, some at home, some in prod and some in test. I need some of it firewalled\private and some of it on the internet. I like all my websites secured by SSL but I'm far to lazy to deal with certificates
2 x Webserver Containers
Small cPU and RAM
Shared storage filesystem(Website content, maybe a SQL DB
Exposed to the world via Cloudflare
Plex
Heaps of storage
Connected back to my home network(ZeroTier)
Gitea
Heaps of storage
Exposed to the world via Cloudflare
Occasional console access
Port 22 directly exposed to internet
TP-Link Omada controller
A number of ports exposed to my LAN
Runs at home
HomeAssistant
USB Device passthrough
Pinned to a specific host
Runs at home
A number of ports exposed through the host
1 port exposed to the world via Cloudflare

Virtual machines
(Hyperstack user) I want to boot a machine with GPU’s, I dont care where it is, I need it to be always available so if it stops, restart it, if the host crashes then respawan it.  i need SSH access on a public IP and i might need one or more ports open. Cloudflare proxy would be ideal for my web applications.
I want to run a VM, It's a pet, so the data is important. If the VM stops, tell me and restart it automatically. I dont need it to have a public IP but i do need to be able to access it securely via SSH. Tailscale, netbord etc would be an acceptable option
I need to deploy a large number of VM's, i want them all on different hosts, some in different datacetners. They are cattle so i dont care about the data, but i do care abaout availability
I'd like to join my dev machine or laptop into the cloud as a deployment location, so I can run containers or VM's etc on this machine with no hosting costs, when i want to scale i can launch on cloud machines. VM's and containers i launch on my machine can seamlessly communciate with VM's and containers hosted in the cloud(THis is fucking amazing!!)
I want to spin up a VM and have it do something for me on boot then adivse me of the results, so when I launch the VMi should eb able to rpovide a bash script or ansible playbook and have it run, the results will be available to me once complete
When I'm managing resource i need a super simple single pain of glass view on the status, if a VM is taking ages to boot, i wanna know why.





Ideas
Datacente operator view
Allow your DC team to see a view on the hardware state
Monitor power consumption of each node
Thermal aware scheduling, use ‘preferential tags’ to help prioritise workload allocation
Advanced health checks, not just ‘availability’ but check the connection to storage, GPU, internet etc
Needs to be infiniband \ Spectrum X frendyl - Does it?
PHysical server location int he DC, floor plan and rack elevation map, show customers on a map
Plugins to hardware providers(Hydrahost, Denver, Iren etc), it will dynamically configure hardware in remote locations on demand, providing a seamless hardware pool
Will detect and report stock levels(If available) in these remote regions
Will shutdown servers that are idle after a period of time to save on power costs
Connect to BMC and shutdown \ bootup as required
For networking, can we simply do VXLAN via the internet using a UDP punch tehcnique and a central coordinator? This would need some firewalls and or encrpytion
Build a feature to ‘catch’ a node when it becomes available to move it into ‘maintenance’
Backfill load and demand onto other providers, by having a scheduling system we can allow users to have visibikity of regiosn that are 'virtual' regions like digital ocean or vultr and the underluying hosts are created on demand
Allow for ‘spot’ style instances - perhaps just set a label on the workload that indicates it is available for instant termination
Rapid spin up and sales though ‘spot’ - onboard stock from a hardware supplier that has a short term availability then add it to the community availability
Workload hosts need some capacity to have an expiration date?
Container in VM for security - If a user wants to be secure, we launch a vm then we run the container\pod inside the VM.. Add’s a memory overhead but limits blast radius
Add firewall blacklists to the region(?) to prevent workloads from routing to DC local subnets.. Could default add the regions ip ranges to the blacklist? The SDN controller will add a block rule(New table?) for any packets destined to that IP range. This would make docker bridge networking viable in some circumstances?!






Full core infra stack
DNS
Cloudflare style proxy(Use cloud flare to begin with?)
Load balancing(Use the docker concept)
VPN (Dial in and Site to Site)
SSH console proxy(Spawns a SSH process on a router namespace on a known random port that will only proxy to the given host
S3 Storage
Create \ Delete S3 creds
Set quota
Get various metrics
Backups (Scheduled or manual) to S3
Snapshots
Send to S3
Launch VM from S3 snap
Images
Public
Private (From Snap)
Private (From upload \ BYO Image)
Custom Public IP’s(BYOIP)






Justification
Kill openstack worker
Strengthen connection between iaas and hyperstack
MASSIVE speed increase
Zero on prem infra requirements
Zero capex on openstack controllers
High availability of control plance because it’s in the ‘cloud’
No firewall or VPN required
No need to listen to rabbit MQ and miss messages, theAPI will callback via HTTP and retry unless it gets a 200
Resize will work
Self updating
No requirement for K8’s argoCD \ Devops teamt o mange complex workers




Zero capex site plan
No management servers - Cloud
No l3 agents- either cloud routed or local vlan
No ceph servers - hyper converged\local raid levels of redundancy
Rapid POC’s
Easily utilise third party systems
‘Region free’ GPU’s - Just deploy this workload to a 5090 please, i dont care where
Instant clusters equivalent, hardware based docker sxm and infinibadn
CPU and GPU laods on GPU nodes when using container alot easier
Dont need master nodes in the same region as the workers if we use VPN!
Karls thing to consume all GPU’s using workload priority
“Today, when a user creates a VM, cluster, or other resource, our API immediately returns a resource ID, implying that the resource exists and can be retrieved using GET /resources/{id}. However, creation may fail at various stages, producing ambiguous or inconsistent outcomes:”






Todo
When creating a container allow binding to a network - Use VXLAN encap on OVS, this is in prep for VM port forarding.
Do we need ot build out a port forwarding database model? Such that we can use it to refresh the namespace container easily? Or we just lookup all the port forward info from eahc associated workload?
When creating a VM - Build a cloud-init CD Image with the relevant metadata - primarily IP address assignment
When creating VM’s - have it add the NIC(s) and bind them to a VXLAN
When creating VM’s - Allow the same port mapping as containers and add an NScontroller on the host




Nexgen Demo
Architecture overview
Deployment time and complexity - single container
Deploy a VM with GPU passtrhough - qemu or cloud-hypervisor
Deploy a container on same network
Both have access to same parallel FS!
No WAN setup required, uses host net(unsafe) or dedicated ovs bridge(safer)
Container is accessible on the internet immediately via cloudflare
VM is accessible on the internet immediately via cloudflare
VM VNC
Container logs and interactive term


Seond host on second site - VM and container launch - Able to comm to eahc other via site-to-site comms


Aggregate equivalent demo
Pod to pod comms, how do we know what IP they have to allow for pod to pod comms
SQL cluster
Proxy front end
Hibernation
Interruptible workloads

Customers
https://cloudification.io/
Vault
Nexgnen
Catalyst
Wafai

Environments
Beta deployment
Server - hvh-node01
Runs cloudflared and hosts these URLs
Dashboard.xcloudify.tech (8080)
Api.xcloudify.tech (5000)
Vnc-console.xcloudify.tech (8085)
Vnc-proxy.xcloudify.tech (6002)
Ws.xcloudify.tech (6001)


Workers
Region 1 - US-FarNorth
Persistent-franklin-0 (Hyperstack) (185.216.21.13)
Brilliant-rutherford (Hyperstack) (185.216.21.90)


Region 2 - Australia1
Caznet-cory-2
Caznet-cory-3
Caznet-cory-4


Region 3 - Gould Creek
	hvh-node02


Region 4 - Norway 1
	Keen-schrodinger (Hyperstack) (
Dev env
Server - Desktop PC
Runs cloudflared and hosts these URLs
api-dev.xcloudify.tech
Workers
Region 1 - 
Lab1
Hvh-node03
hvh-node04


Region 2 - 
DesktopPC
hvh-node05


